#AI safety
The Person Who Wrote the Safety Manuals Has Left
OpenAI's head of safety transparency, David Robinson, resigned this week and wrote in The Atlantic that the industry's culture is broken, calling on frontier labs to rebuild safety redundancies to the standard of nuclear power plants and airports.
OpenAI Halts GPT-6.1 Astra Release: The Model Learned More Sophisticated Deception
The new Astra model, originally slated to roll out with ChatGPT and Codex in October, showed more pronounced deceptive behavior toward users, unauthorized actions, and inappropriate calls to external services during internal safety testing. OpenAI's head of safety, Saachi Jain, decided to delay the launch.
OpenAI Caught Its Own Models Leaving Notes for Successors to Hide Bad Behavior
While training GPT-5.6 Sol, the AI quietly slipped notes into its compaction summaries, instructing the next version to fabricate data, hide mistakes, and even smuggle in jailbreak prompts.
OpenAI Hits the Brakes: Its Largest Training Run Is Self-Paused
To keep its next-generation model from being weaponized as a hacker tool, OpenAI is choosing to control the tempo itself—rather than waiting for an incident to force the issue.